Explanation-based Learning and Reinforcement Learning: a Uniied View

نویسنده

  • Andrew Barto
چکیده

In speedup learning problems where full descriptions of operators are known both explanation based learning EBL and reinforcement learning RL methods can be applied This paper shows that both methods involve fundamentally the same process of propagating information backward from the goal toward the starting state Most RL methods perform this propagation on a state by state basis while EBL methods compute the weakest preconditions of operators and hence perform this propagation on a region by region basis Barto Bradtke and Singh have observed that many algorithms for reinforcement learning can be viewed as asynchronous dynamic programming Based on this observation this paper shows how to develop dynamic programming versions of EBL which we call region based dynamic programming or Explanation Based Reinforcement Learning EBRL The paper compares batch and online versions of EBRL to batch and online versions of point based dynamic programming and to standard EBL The results show that region based dynamic programming combines the strengths of EBL fast learning and the ability to scale to large state spaces with the strengths of reinforcement learning algorithms learning of optimal policies Results are shown in chess endgames and in synthetic maze tasks

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

Explanation-based Learning and Reinforcement Learning: a Uniied View

In speedup-learning problems, where full descriptions of operators are always known, both explanation-based learning (EBL) and reinforcement learning (RL) can be applied. This paper shows that both methods involve fundamentally the same process of propagating information backward from the goal toward the starting state. RL performs this propagation on a state-by-state basis, while EBL computes ...

متن کامل

Survey of effective factors on learning motivation of clinical students and suggesting the appropriate methods for reinforcement the learning motivation from the viewpoints of nursing and midwifery faculty, Tabriz University of Medical Sciences 2002.

Introduction. Motives are the powerful force in process of education– learning, so that the richest and best training plans and structured education are not effective if the lack of motivation existed. In spite of the fact that the success of teacher depends on the learning motivation of students, then it is necessary for teachers to know the effective methods for motivating the students and t...

متن کامل

Explanation of Harry Broudy’s View with Respect to Aesthetic Education and its Link to Education via Pedagogical Theater

Pedagogical theater with a continuous process and with an emphasis on the simple learning of various concepts and lessons assists to grow and thus enhance individual and group behavior in society. The process of performing the exercises encourages the talent and creativity of the participants in learning and ensures their active participation. This study aims to establish a link between the ide...

متن کامل

Dynamic Obstacle Avoidance by Distributed Algorithm based on Reinforcement Learning (RESEARCH NOTE)

In this paper we focus on the application of reinforcement learning to obstacle avoidance in dynamic Environments in wireless sensor networks. A distributed algorithm based on reinforcement learning is developed for sensor networks to guide mobile robot through the dynamic obstacles. The sensor network models the danger of the area under coverage as obstacles, and has the property of adoption o...

متن کامل

Reinforcement Learning Based PID Control of Wind Energy Conversion Systems

In this paper an adaptive PID controller for Wind Energy Conversion Systems (WECS) has been developed. Theadaptation technique applied to this controller is based on Reinforcement Learning (RL) theory. Nonlinearcharacteristics of wind variations as plant input, wind turbine structure and generator operational behaviordemand for high quality adaptive controller to ensure both robust stability an...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

عنوان ژورنال:

دوره   شماره 

صفحات  -

تاریخ انتشار 1995